Papers with academic research
Advances in Debating Technologies: Building AI That Can Debate Humans (2021.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial focuses on Debating Technologies, a sub-field of computational argumentation defined as "computational technologies developed directly to enhance, support, and engage with human debating" the tutorial provides a holistic view of a debated system, and discusses practical applications and future challenges of debation technologies. |
| Approach: | They present a tutorial on Debating Technologies, a sub-field of computational argumentation . they introduce Project Debater, which is the first AI system to debate human experts . |
| Outcome: | The project Debater is the first AI system to debate human experts on complex topics. |
Autonomous Machine Learning-Based Peer Reviewer Selection System (2025.coling-demos)
Copied to clipboard
| Challenge: | Existing systems that match papers with experts are inefficient and often require long turnaround times. |
| Approach: | They propose an autonomous peer reviewer selection system that employs the natural language processing model to match submitted papers with expert reviewers independently of traditional journals and conferences. |
| Outcome: | The proposed system performs faster and smaller than current models while being more scalable. |
EduPulse: A Practical LLM-Enhanced Opinion Mining System for Vietnamese Student Feedback in Educational Platforms (2026.eacl-industry)
Copied to clipboard
| Challenge: | EduPulse is a system designed specifically to analyze student feedback in Vietnamese. |
| Approach: | They propose a system that analyzes student feedback in Vietnamese to improve opinion mining. |
| Outcome: | The proposed system performs four opinion analysis tasks in Vietnamese . it is scalable and maintainable, and it is cost-effective, the authors show . |
CLLE: A Benchmark for Continual Language Learning Evaluation in Multilingual Machine Translation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks for Continual Language Learning (CLL) are limited due to the complexity of the task and the lack of unified benchmarks. |
| Approach: | They propose a Continual Language Learning Evaluation benchmark CLLE in multilingual translation. |
| Outcome: | The proposed method is effective when compared with other strong benchmarks. |
The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions (2023.emnlp-main)
Copied to clipboard
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han
| Challenge: | Recent advances in large language models (LLMs) have produced models that exhibit remarkable performance across a variety of NLP tasks. |
| Approach: | They analyze a large-scale collection of user-GPT conversations to identify a significant gap between academic research in NLP and the needs of real-world NLP applications. |
| Outcome: | The proposed model outperforms existing models in a large-scale collection of user-GPT conversations and identifies a significant gap between the tasks that users frequently request from LLMs and the tasks commonly studied in academic research. |
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains (2025.findings-naacl)
Copied to clipboard
| Challenge: | a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech . |
| Approach: | They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection . |
| Outcome: | The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture . |
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based English Information Retrieval Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | ad-hoc information retrieval methods usually require large amounts of annotated data to be effective. |
| Approach: | They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia. |
| Outcome: | The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs. |
On the Limitations of Simulating Active Learning (2023.findings-acl)
Copied to clipboard
| Challenge: | Active learning (AL) is a human-and-model-in-the-loop paradigm that iteratively selects informative unlabeled data for human annotation. |
| Approach: | They propose to simulate active learning by using an already labeled dataset as the pool of unlabeled data. |
| Outcome: | The proposed model-in-the-loop paradigm can be used to perform experiments with human annotations on-the fly. |
ResearchArena: Benchmarking Large Language Models’ Ability to Collect and Organize Information as Research Agents (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models excel across many natural language processing tasks but face challenges in domain-specific, analytical tasks such as conducting research surveys. |
| Approach: | They propose a benchmark to evaluate LLMs' capabilities in conducting research surveys. |
| Outcome: | The proposed benchmark is designed to evaluate LLMs' capabilities in conducting research surveys. |
BabyCloud, a Technological Platform for Parents and Researchers (L18-1)
Copied to clipboard
Xuân-Nga Cao, Cyrille Dakhlia, Patricia Del Carmen, Mohamed-Amine Jaouani, Malik Ould-Arbi, Emmanuel Dupoux
| Challenge: | a platform for capturing, storing and analyzing day-long audio recordings and photos of children's linguistic environments is proposed . the proposed platform connects families and academics, with strong innovation potential for each type of users. |
| Approach: | They propose a platform for capturing, storing and analyzing audio recordings and photos of children's linguistic environments. |
| Outcome: | The proposed platform connects families and academics with strong innovation potential for each type of users. |
Is GPT-4V (ision) All You Need for Automating Academic Data Visualization? Exploring Vision-Language Models’ Capability in Reproducing Academic Charts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Using Vision-Language Models (VLMs) for data visualizations requires significant time and expertise in both data management and graphic design. |
| Approach: | They propose a dataset comprising 2525 high-resolution data visualization figures with captions from AI conferences, extracted directly from source codes. |
| Outcome: | The proposed model outperforms open-source models in reproducing complex charts while using Chain-of-Thought prompting. |
“A good pun is its own reword”: Can Large Language Models Understand Puns? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the understanding of puns in large language models (LLMs) have not explored the use of pun in creative writing and humor creation. |
| Approach: | They propose to use pun recognition, explanation and generation tasks to evaluate the capabilities of large language models (LLMs) they adopt automated evaluation metrics from prior research and introduce new evaluation methods and metrics that align more closely with human cognition. |
| Outcome: | The proposed methods align more closely with human cognition than previous evaluation metrics. |
Corpus Building and Evaluation of Aspect-based Opinion Summaries from Tweets in Spanish (L18-1)
Copied to clipboard
Daniel Peñaloza, Rodrigo López, Juanjosé Tenorio, Héctor Gómez, Arturo Oncevay-Marcos, Marco A. Sobrevilla Cabezudo
| Challenge: | a corpus of Spanish extractive and abstractive summaries of opinions is presented . the goal is to analyze the summary content and to show how different they are written . |
| Approach: | They present a corpus of Spanish extractive and abstractive summaries of opinions . they analyze the summary agreement between them and their aspect coverage and sentiment orientation . |
| Outcome: | The presented corpus of Spanish extractive and abstractive summaries is a reference for academic research. |
PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs (2023.emnlp-main)
Copied to clipboard
Rahul Goel, Waleed Ammar, Aditya Gupta, Siddharth Vashishtha, Motoki Sano, Faiz Surani, Max Chang, HyunJeong Choe, David Greene, Chuan He, Rattima Nitisaroj, Anna Trukhina, Shachi Paul, Pararth Shah, Rushin Shah, Zhou Yu
| Challenge: | PRESTO dataset contains 550K contextual multilingual conversations between humans and virtual assistants. |
| Approach: | They propose to use a dataset of 550K contextual multilingual conversations between humans and virtual assistants to study some of the more challenging aspects of parsing realistic conversations. |
| Outcome: | The dataset contains 550K contextual conversations between humans and virtual assistants. |
ReviewEval: An Evaluation Framework for AI-Generated Reviews (2025.findings-emnlp)
Copied to clipboard
| Challenge: | escalating volume of academic research necessitates innovative approaches to peer review . authors propose reviewEval, ReviewAgent and ReviewEval to improve on existing reviews . |
| Approach: | They propose a framework for AI-generated reviews that measures alignment with human assessments . they propose 'reviewAgent' that iteratively optimizes its intermediate outputs and external improvement loops . |
| Outcome: | The proposed framework improves actionable insights and analytical depth by 6.78% and 47.62% over baselines and expert reviews. |